Repository navigation
ROCm: compress the embedded GPU code; upgrade llama.cpp b11247 → b11256 - #470
Merged
Merged
Conversation
Nine upstream commits, one reviewable step (22 KiB diff). No project source change: #29632 moves several tools/examples to llama_backend_init(), which jllama, TTS and the trainer already call; fs_write_atomic() (#29642) is unused here; the scheduler graph_inputs pass (#29634) and the faster GGUF duplicate checks (#29598) are internal. server-schema.cpp and server-task.cpp are untouched, so the request/response contract is unchanged. All nine patches apply unchanged and are all still needed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
The Windows ROCm jllama.dll was ~1 GB: every HIP translation unit embeds one code object per GPU target, i.e. the whole kernel set once per architecture, and clang stores those bundles uncompressed by default. The jar hid it (234 MB zipped), but LlamaLoader extracts the library to the temp dir on every start, and in the all-backends fat jar ROCm is tried right after CUDA - on nearly every machine without an NVIDIA card. ggml-hip is now compiled with --offload-compress, so each bundle is stored zstd-compressed (CCOB) and inflated by the HIP runtime at module load. The option is scoped to the ggml-hip target and to its source language (HIP on Linux, CXX on Windows, where upstream compiles HIP as C++). Both ROCm jobs run .github/verify-hip-offload-compressed.py after the build: it prints the library size and bundle counts (also into the job summary) and fails on any uncompressed bundle, or on none compressed, so a toolchain or upstream change cannot quietly bring the 1 GB library back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
bernardladenthin
had a problem deploying
to
startgate
September 29, 2026 15:56 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
September 29, 2026 15:56 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
maven-central
September 29, 2026 15:56 — with
GitHub Actions
Failure
CLAUDE.md said our target lists add gfx900/gfx906/gfx90c/gfx1153 to upstream's. That holds for Linux only: upstream's windows-rocm list already carries gfx1153 (checked at b11247 and b11256), so the Windows extras are just gfx900/gfx906/gfx90c. Same correction in the Windows ROCm job's comment. No target list changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
bernardladenthin
had a problem deploying
to
maven-central
September 29, 2026 15:57 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
startgate
September 29, 2026 15:57 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
September 29, 2026 15:57 — with
GitHub Actions
Failure
|
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
--offload-compress). The Windows ROCmjllama.dllinllama-*-all-windows-x86-64-jar-with-dependencies.jarwas ~1 GB. Each HIP translation unit embeds one code object per GPU target (the whole kernel set, once for each of the 23 architectures), and clang stores those bundles uncompressed by default.LlamaLoaderextracts the library to the temp dir on every start. In the all-backends fat jar, ROCm is tried right after CUDA, so this happens on nearly every machine without an NVIDIA card.ggml-hipis now compiled with--offload-compress, scoped to that target and its source language (HIP on Linux, CXX on Windows). The HIP runtime decompresses each bundle when the module loads..github/verify-hip-offload-compressed.pyafter the build.__CLANG_OFFLOAD_BUNDLE__), or if there is no compressed one at all, so losing the flag reds the job instead of shipping the 1 GB library again.llama_backend_init(), which jllama, TTS and the trainer already call.fs_write_atomic()(#29642) is unused here; the schedulergraph_inputspass (#29634) and faster GGUF duplicate checks (#29598) are internal.server-schema.cppandserver-task.cppare untouched, so the request/response contract is unchanged.release.ymlchange, so the CUDA/ROCm/OpenVINO pins stay.docs/history/llama-cpp-breaking-changes.md.Test plan
jllama+jllama_testsucceedsctest: 590/590 C++ tests passNativeLibraryLoadSmokeTest4/4 (linked build-info matches theb11256pin)--offload-compressand the resulting DLL/.so size. The two ROCm jobs' new "Verify the GPU code is compressed" step and its job-summary table show both.Related issues / PRs
Refs #468 (b11237 → b11247)
Checklist
🤖 Generated with Claude Code
https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
Generated by Claude Code